BMC Genomics
Top medRxiv preprints most likely to be published in this journal, ranked by match strength.
Show abstract
MotivationFanconi anemia (FA) is a rare disease mainly caused by biallelic pathogenic variants, including structural variants such as large deletions and insertions in FA genes. Currently, variant detection is based on short-read sequencing and probe-based approaches. However, determining the exact genomic breakpoint or achieving allelic discrimination remains challenging. Nanopore-based long-read sequencing enables a comprehensive detection of FA variants, but a unified bioinformatic analysis p...
Show abstract
Accurate classification of BRCA1 and BRCA2 variants is essential for cancer risk assessment and therapy selection, yet over one-third remain variants of uncertain significance (VUS). Here, using 120,660 real-world cancer genomic profiles with BRCA1 or BRCA2 variants from a >800,000-sample cohort, we develop machine learning models that predict pathogenicity using clinical and tumor-derived features, including a pan-cancer homologous recombination deficiency signature, co-mutated genes, zygosity,...
Show abstract
BackgroundKlebsiella pneumoniae is a common cause of neonatal sepsis in Africa, and is frequently hospital acquired. We recently reported an outbreak of multidrug-resistant K. pneumoniae sepsis amongst neonates at a rural hospital in The Gambia, West Africa, involving 57 cases and case fatality of 60%. Here we undertook a retrospective pathogen genomic epidemiology study of clinical and environmental K. pneumoniae isolated during the outbreak, to identify the outbreak strain, refine the epidemic...
Show abstract
Rare Mendelian disorders affect 300-400 million people globally. Although genetic testing has become widely adopted, gene-specific evidence for tailored variant interpretation remains scattered across resources. We present Gene Portals, a framework for gene-centered multimodal knowledge bases that co-localize expert-harmonized clinical data, functional assays, population variation, structural annotations and gene-specific ACMG/AMP specifications within a single resource. A modular interface inte...
Show abstract
Tumour typing from whole-genome sequencing is increasingly accurate, yet molecular subtyping from somatic variants remains challenging because of tumour heterogeneity and inconsistent clinical annotations. Here, we present Mutation-Attention Dual-Task (MuAt2), a Transformer model that jointly classifies histological tumour types and subtypes directly from somatic single-nucleotide variants, indels and structural variants. MuAt2 leverages encoders pre-trained on 2,587 pan-cancer whole genomes, an...
Show abstract
BackgroundA coronary artery calcium (CAC) score of 0 is widely considered to indicate low short- to intermediate-term risk for coronary artery disease (CAD) and is frequently used to defer lipid-lowering therapy. However, a subset of individuals with CAC=0 still experience events, highlighting residual risk not captured by imaging alone. Polygenic risk scores (PRS) quantify lifelong inherited susceptibility, but conventional approaches rely on predefined ancestry labels despite human genetic div...
Show abstract
Type 2 diabetes (T2D) affects 11.1% of the global population, underscoring the need for biomarkers that inform treatment response and glycemic outcomes. We evaluated the association between the FTO variant rs9939609-A and glycemic control in a Mexican population. A total of 174 individuals living with T2D from Merida and Sisal, Yucatan, were included, of whom 85% were receiving oral hypoglycemic agents as main treatment. Glycemic control was defined cross-sectionally as good ([≤]130 mg/dL, n=...
Show abstract
We report a previously undescribed genotypic configuration identified in twins with HNRNPU-related neurodevelopmental disorder. Both twins have two closely spaced mosaic variants on the same allele that never co-occur on any single DNA molecule, resulting in three distinct cell lineages within each individual. We define this genotypic configuration as clustered monoallelic mosaicism (cMoMa). Recognizing the extreme improbability of such a configuration, we systematically explore two potential me...
Show abstract
Chromosome 5p15.33 harbors several independent association signals which demonstrate antagonistic pleiotropy across cancer types, with causal mechanisms largely unresolved. To identify functional variants and enhancer elements at this locus, we performed statistical fine-mapping followed by massively parallel reporter assays (MPRA) and proliferation based CRISPRi screens. This approach identified eight multi-cancer functional variants (MCFVs) across three GWAS signals. Targeting rs421629 (part o...
Show abstract
PurposeQuantitative metrics obtained from retinal fundus images (such as vessel length, tortuosity and other scale-dependent measures) are increasingly used as potential biomarkers for systemic diseases, including cardio- and neurovascular conditions. However, with the increasing prevalence of myopia and related axial growth, this study aims to evaluate if axial length scaling significantly alters the overall distributions of the inferred biomarkers when compared to biomarker data obtained witho...
Show abstract
PurposeTo investigate the effects of morning and evening narrowband blue light exposure on axial length, and to examine the short-term effect of morning blue light combined with myopic defocus on axial length. MethodsFor objective 1, 18 individuals underwent 60 minutes of narrowband blue light exposure (460nm) in the morning (9:00-11:00AM) and evening (5:00-7:00PM) of the same day. The axial length values were normalized to the average of the morning and evening axial length values. For objecti...
Show abstract
Nocturnal glucose regulation is modulated by autonomic and circadian mechanisms, yet their dynamic interplay in apparently healthy, free-living populations remains poorly studied. Here, we assessed 227,860 nights of concurrent sleep data from Ultrahuman AIR ring and M1 continuous glucose monitoring (CGM) system across 5849 adults globally to examine nocturnal cardio-metabolic coupling. We found that higher sleep consistency was inversely associated with glucose variability, and vice versa. Unsup...
Show abstract
ObjectivesTo identify unique echocardiographic signatures associated with TTR+ carrier status preceding onset of cardiac amyloidosis. BackgroundCarrier status for the most common pathogenic TTR variant in the United States, Val142Ile (V142I), found in 4% of African Americans (AA) and 1% of Hispanic/Latino (H/L) individuals, confers a 40-60% lifetime risk of developing variant transthyretin amyloidosis (ATTRv), including cardiac amyloidosis (CA) and heart failure (HF). Myocardial amyloid deposit...
Show abstract
Drug-induced liver injury (DILI) is an acute inflammatory liver disease caused not only by prescription and over-the-counter medications but also by health foods and dietary supplements. Typically, DILI patients recover once the causative substance is identified and discontinued. In contrast, autoimmune hepatitis (AIH) results from the immune-mediated destruction of hepatocytes due to a breakdown of self-tolerance mechanisms. Patients presenting with acute-onset AIH often lack characteristic cli...
Show abstract
BackgroundHypertension affects over 30% of adults and is the leading risk factor for cardiovascular disease. It often presents without obvious symptoms, meaning that, although effective therapies exist, hypertension remains widely undiagnosed and insufficiently treated. Genomics-based prediction methods have shown only modest benefits for these disorders, but proteomic markers have demonstrated potential for greater predictive and clinical value. MethodsWe applied a novel machine-learning based...
Show abstract
Wastewater monitoring enables non-invasive, population-scale tracking of community infections independent of healthcare-seeking behavior and clinical diagnosis. Metagenomic sequencing extends this capability by enabling broad, pathogen-agnostic detection, genomic characterization, and identification of novel or unexpected threats. Here, we present data from CASPER (the Coalition for Agnostic Sequencing of Pathogens from Environmental Reservoirs), a U.S.-based wastewater metagenomic sequencing ne...
Show abstract
IntroductionGenome-wide association studies (GWAS) for kidney function have mainly focused on creatinine-based glomerular filtration rate (eGFRcrea), which is affected by variation in muscle mass. Moreover, the genetic basis of the sexual dimorphism of chronic kidney disease is underexplored. MethodsWe performed a GWA meta-analysis for creatinine clearance (CrCl), a muscle mass-independent kidney function phenotype, in 58,976 individuals of European descent from the Lifelines Cohort Study. Res...
Show abstract
STUDY QUESTIONAre pathogenic variants in Homeodomain-interacting protein kinase (HIPK4) associated with sperm head abnormalities causing male infertility? SUMMARY ANSWERHIPK4 is a novel candidate gene associated with sperm head defects and human male infertility. WHAT IS KNOWN ALREADYNumerous genes causing male infertility due to Multiple Morphological Abnormalities of the sperm flagella (MMAF) have been described but the genetic basis of sperm head defects is less well understood. STUDY DESI...
Show abstract
Background: Pressure volume (PV) loop analysis remains the gold standard for assessing the intrinsic global diastolic properties of the left ventricle (LV). Traditional fitting techniques rely on local, phase-constrained fittings and are limited due to their sensitivity to noise, landmark selection, violation of assumptions, and non-convergence. Objective: To develop and validate DIAPINN, a physics-informed neural network (PINN) framework capable of calculating intrinsic diastolic properties of ...
Show abstract
BackgroundThe relationship between hip osteoarthritis (hip OA) and Alzheimers disease (AD) presents a critical paradox within the emerging "bone-brain axis": widespread phenotypic comorbidity sharply contradicts evolutionary theories of biological antagonism. This study integrates longitudinal and multi-omic analyses to determine whether this clinical overlap masks an underlying genetic neuroprotection. MethodsWe analyzed longitudinal phenotypic data from 261,767 UK Biobank participants using C...